[Perf] Add opt-in DFlash2 schedules and numerical audits - #556
Merged
yangzhuxinyzx merged 40 commits intoSep 9, 2026
Merged
Conversation
…al audits Record accepted-slot provenance and reject incomplete four-rank comparisons. Localize the observed A/A drift to autotuned first-layer Gemma RMSNorm reduction order. Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Pin reduction extent and warp count for FP16 input and residual norms. Preserve masked-square and FP32 residual materialization boundaries to match the recorded 8192-element Inductor path. Keep disabled pending natural-output and complete-round promotion gates. Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Read the Qwen3.5 projection row stride directly and retain FP32 beta. Require per-forward route evidence and add real-state replay and runtime-bridge parity coverage. Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
4 tasks
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Honor explicit prefill metadata in the no-active-spec GDN branch so a one-token initial prefill does not consume recycled conv/SSM state. Preserve real decode, legacy no-flag, and non-speculative behavior. Add CPU regressions for singleton routing, cached initial-state flags, graph metadata settings, and recycled finite/NaN convolution state. Related: vllm-project/vllm#51565 (narrow 1Cat fork adaptation). Signed-off-by: Zhaochengggg <87113558+zhaochengggg@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Add the warp memory barrier before online-softmax state publication and retain a short q8 racecheck fixture. Record the native kernel, fixed-prefix, and complete-round candidate results without enabling unadmitted routes. Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Preserve the repaired per-head arithmetic while building one-group and three-group private candidates with complete source manifests. Record the rebuilt numerical gate and the bounded QPN2 scheduling experiments. Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Record the actual four-rank fixed-prefix attention gate and unprofiled pair. Add isolated two-channel publication experiments; both serialized and overlapped chunks regress and remain disabled. Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Add an explicit experimental installer that preserves the CPU predicate and defers cache commit. Record four-rank natural shadow parity and qualify the first timing pair. Add the actual-weight FP16 layout error screen and retain rejected broad variants. Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Retain post-reboot context overlap proof and the rejected QPN2 and native FlashInfer fragment screens. Keep changed draft trajectories unadmitted and gate the GDN schedule to captured TP4 q8 with FP32 state. Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
…ash2-15ms-20260907-161715 Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com> Assisted-by: Codex
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
…flash2-15ms-20260907-161715 Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
yangzhuxinyzx
marked this pull request as ready for review
September 9, 2026 12:50
Assisted-by: Codex Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
DFlash2 TP4/B1/q8 verification spends substantial time in projection, state layout, attention and scheduling. This change adds independently selectable schedules and reproducible operator/model audits for the current approximately 16-ms complete-round path. The common installer selects GDN BV2, context/probe scheduling, grouped E4M3 attention and sparse selection through actual tensor/shape guards, independently of target quantization names. A separate hashed QPN2 manifest selects NVFP4 cap64 projections and TP4 publication. Unsupported shapes retain their existing operators.
The PR also contains the grouped-attention warp synchronization repair and reuses dependency #563's singleton-prefill metadata fix. It extends existing PR #556 rather than creating duplicate optimization or prefill-fix PRs. No experimental route is enabled by default; source integration does not certify unfinished runtime admission. Sub-15-ms complete rounds have not been achieved.
Test Plan
Freeze source, native libraries, model weights, sampling and physical GPUs 4–7. Preserve TP4/B1/q8, E4M3 target KV, FP32 logits/state, FP16 draft transport, T1/k20/p.95/xhigh and natural EOS. CUDA 12.8, Torch 2.10.0+cu128 and Python 3.12.13 are pinned. The server has
--max-model-len 262144; each natural request usesmax_tokens=262144-actual_prompt_tokensverified with the server tokenizer. Input is not truncated. One-token warmups and bounded operator diagnostics are excluded from natural-generation scoring.The frozen corpus contains 148 prompts: 32 each GSM8K, MATH500, HumanEval and MBPP, 16 LiveCodeBench v6 and four JSON/tool cases, with seeds 0/1/2. Independent startups use five warmups and five measured requests per speed fixture/arm. Whole-stack all-off controls, all-layer repeatability, public-entry parity, FP8 and long-prefix boundaries are separate gates. These are dataset subsets, not full-dataset benchmark claims.
Test Result
result > 0semantics; eight boundary checks pass. Original assertions and EvalPlus eligible subsets have separate denominators.CUDA_VISIBLE_DEVICES='' .venv/bin/python -m pytest tests/v1/attention/test_gdn_metadata_builder.py --confcutdir=tests/v1/attention -q -k 'singleton_prefill or gdn_build_classification or common_gdn_metadata_matches or mixed_decode_stays_decode_fastpath or full_cuda_graph_decode_padding_uses_pad_slot'(32 passed, 30 deselected).pre-commit run --files docs/design/sm70_dflash2_acceptance_20260909.md. Rejected arithmetic or slower scheduling experiments remain disabled and documented.Integration and remaining validation
Main base
b6d91d61ff030fe325dc69eb8372a5c72d0374bdis already included. The running evaluation checkout remains frozen ata7cc5ae305149d7a9ffdf42fb224dff34e5606aa; integration and documentation work do not change its source or native libraries. Reproduction commands, hashes, numerical evidence, long-generation observations and artifact locations are retained indocs/design/sm70_dflash2_acceptance_20260909.md,docs/design/sm70_quasar_dflash2_15ms.mdanddocs/design/sm70_quasar_dflash2_resource_audit_20260909.md.This integration preserves opt-in selection. Complete 256K-capacity multi-seed scoring, independent starts, whole-stack quality/acceptance, all-layer repeatability, public-installer parity, real FP8 and long-prefix validation remain open before any default promotion. The previous 4.33% repeat-start TV discrepancy is not an allowed error margin. No serving API, weights, model format, sampling semantics or context capacity is changed by integration.
AI assistance was used for implementation, diagnostics and reporting.
Assisted-by: Codex